Papers with reward-model-based RLHF

    1 papers
    Filtered Direct Preference Optimization (2024.emnlp-main)

    Copied to clipboard

    Challenge: Existing studies on the impact of RLHF on text quality have focused on reward-model-free RL.
    Approach: They propose an extension of direct preference optimization to improve model performance by analyzing the quality of the preference dataset.
    Outcome: The proposed method improves the performance of models optimized with DPO over those optimized with reward-model-based RLHF.

    What is GenGO?

    GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

    Information

    About
    Limitations